Papers with human annotation
Copied to clipboard
| Challenge: | Existing methods to enable skill routing do not scale in terms of the number of skills and skill on-boarding. |
| Approach: | They propose a model-based approach to enable natural conversation by allowing frequent policy updates . they propose an annotation-based system, rule-based model, and bandit-based learning . |
| Outcome: | The proposed method is scalable and cost-effective, the authors show . they show that it can improve the user experience without abrupt policy changes . |
Copied to clipboard
| Challenge: | NER is a fundamental problem for medical text mining because of the difference of specialties and cost of human annotation. |
| Approach: | They propose a label-aware double transfer learning framework for medical NER from electronic medical records. |
| Outcome: | The proposed framework improves accuracy over strong baselines on 12 cross-specialty NER tasks. |
Copied to clipboard
| Challenge: | Various attempts to correct noisy data in the construction process have been made, but human annotation is expensive and time-consuming. |
| Approach: | They propose to use large language models for data annotation to imitate human annotation and classify unrelated documents from a multi-document summarization task. |
| Outcome: | The proposed method imitates human annotation and classifies unrelated documents from the Multi-News dataset. |
Copied to clipboard
| Challenge: | Variation in human annotation (i.e., disagreements) is common in NLP, but it is unclear whether it is possible to model this variation in LLMs. |
| Approach: | They evaluate the influence of different reasoning settings on LLM disagreement modeling . RLVR-style reasoning degrades performance in disagreement modeling, they find . |
| Outcome: | The proposed reasoning settings improve LLM disagreement modeling, while RLVR-style reasoning degrades it. |
Copied to clipboard
| Challenge: | Recent advances in NLP have enabled the use of text-to-text annotation without providing training samples. |
| Approach: | They propose a text-to-text interface for automatic annotation using written guidelines without providing training samples. |
| Outcome: | The proposed approach is comparable with the fine-tuned BERT but without any training data. |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucinations in model-generated texts fail to pinpoint errors. |
| Approach: | They propose a formalism for localizing factual inconsistencies in attributable text generation . they propose to decompose the generated text into simple question-answer pairs . |
| Outcome: | The proposed method achieves substantial inter-annotator agreement while achieving a substantial consistency score. |
Copied to clipboard
| Challenge: | Question Answering (QA) is a growing area of research . state-of-the-art QA models struggle on out-of domain documents without fine-tuning . |
| Approach: | They propose a pipeline for validating and training QA data and an interface for human annotation. |
| Outcome: | The proposed pipeline improves QA performance on domain-specific datasets while preserving the accuracy of the model. |
Copied to clipboard
| Challenge: | In contrast, adversarial attacks can cause model errors by modifying inputs, such as the universal triggers attack. |
| Approach: | They propose a data poisoning attack that allows an adversary to control model predictions whenever a desired trigger phrase is present in the input. |
| Outcome: | The proposed attack can cause model errors by modifying inputs, but it can also cause extra human annotation. |
Copied to clipboard
| Challenge: | Current efforts focus on textual claims sourced mainly from Twitter . lack of automated control measures and reliance on human annotation increase noise risk . |
| Approach: | They propose to use a framework to integrate data annotation to mitigate misinformation . they propose to include fact-checks alongside the corresponding claims made by politicians . |
| Outcome: | The proposed dataset will include fact-checks alongside the corresponding claims made by politicians. |
Copied to clipboard
| Challenge: | Recent advances in LLMs enable sophisticated user simulations that can replace traditional rule-based evaluations. |
| Approach: | They propose a persona-driven approach to conversational agent evaluation using Large Language Models (LLMs) they introduce a dataset of customer personas, which are then used to configure a single LLM-based user simulator. |
| Outcome: | The proposed model emulates nuanced customer roles and can implement cross-selling strategies with minimal impact on customer satisfaction, varying by customer type. |
Copied to clipboard
| Challenge: | Existing models for question answering are limited in the availability of labeled data. |
| Approach: | They propose a hierarchical conditional variational autoencoder for generating QA pairs given unstructured texts as contexts while maximizing mutual information between generated QA pair to ensure consistency. |
| Outcome: | The proposed framework achieves impressive performance gains over baseline models on both tasks, using only a fraction of data for training. |
Copied to clipboard
| Challenge: | Recent studies have raised concerns regarding the hallucination and flaws in their reasoning process. |
| Approach: | They propose a framework to learn planning-based reasoning through Direct Preference Optimization on collected trajectories, which are ranked according to synthesized process rewards. |
| Outcome: | The proposed model surpasses GPT-3.5-Turbo on logical reasoning benchmarks on a set of logically-based reasoning tasks. |
Copied to clipboard
| Challenge: | Existing evaluation methods for long-context large language models are overly simplistic and require extensive human annotations. |
| Approach: | They propose an automatic toolkit to create realistic evaluation benchmarks . they use a document-grounded benchmark to generate question-answer pairs . |
| Outcome: | The proposed toolkit provides a way to create realistic evaluation benchmarks and visualize performance metrics of evaluated models. |
Copied to clipboard
| Challenge: | Existing datasets rely on human demonstrations, limiting scalability. |
| Approach: | They propose a scalable data synthesis pipeline that transforms noisy rollouts into reliable supervision without human annotation. |
| Outcome: | The proposed pipeline transforms noisy rollouts into reliable supervision without human annotation. |
Copied to clipboard
| Challenge: | Recent advances in conversational information seeking (CIS) suggest a remedy for the lack of interactive clarification when people face unfamiliar domains. |
| Approach: | They propose a fully autonomous conversational information-seeking agent that couples large language models with a set of domain-specific tools to provide product demand clarification. |
| Outcome: | The proposed agent can iterate over 2,000 automatically generated sessions and score high on real-world evaluations without human annotation. |
Copied to clipboard
| Challenge: | Language of study extraction is an aspect of computational linguistics papers that is useful for analyses of trends and diversity in computational linguists. |
| Approach: | They propose to benchmark and evaluate automated language of study extraction from computational linguistics papers. |
| Outcome: | The proposed language extraction benchmarks show that they can extract languages from papers with accuracy without high computational costs. |
Copied to clipboard
| Challenge: | Documents that are image-based are difficult to extract because of document variability. |
| Approach: | They propose a human-in-the-spiral assistive document annotation platform to extract structured data from document collections. |
| Outcome: | The proposed framework reduces annotation time by at least 41% while showing consistent performance gains over three iterations. |
Copied to clipboard
| Challenge: | Empirical natural language processing (NLP) systems involve interoperation among multiple components . a wealth of NLP toolkits exist ( 4), such as spaCy, DKPro, CoreNLP. |
| Approach: | They propose a unified open-source framework that supports fast development of NLP workflows . framework includes processors for NLP tasks, visualization, and annotation . |
| Outcome: | The framework offers processors for NLP tasks, visualization, and annotation, and is extensible . it is delivered through two modularized yet integratable open-source projects, Forte and Stave . |
Copied to clipboard
| Challenge: | Using FreebaseQA, we can generate over 54K matches from about 28K unique questions with minimal cost. |
| Approach: | They propose a data set for open-domain factoid question answering tasks over structured knowledge bases, like Freebase, using a combination of trivia-type question-answer pairs and subject-predicate-object triples. |
| Outcome: | The proposed data set generates 54K matches from 28K unique questions with minimal cost. |
Copied to clipboard
| Challenge: | Existing methods to replace human annotation are expensive and limited. |
| Approach: | They investigate the use of synthetic data in Fact Verification and Evidence-based Question Answering by replacing human-generated data with synthetic points on eight diverse datasets. |
| Outcome: | The proposed method shows promise but performance declines when replacing up to 90% of training data with synthetic data are severe . the proposed method can be used to improve models trained on purely synthetic data by including as few as 125 human-generated data points. |
Copied to clipboard
| Challenge: | Existing methods to solve Math Word Problems rely on human annotation . empirical results suggest that our method universally improves the performance on single-unknown and multiple-un unknown benchmarks. |
| Approach: | They propose a controlled equation generation solver by leveraging a set of control codes to guide the model to consider certain reasoning logic and decode the corresponding equations expressions transformed from the human reference. |
| Outcome: | The proposed method improves performance on single-unknown and multiple-un unknown benchmarks with 13.2% accuracy on the challenging multiple-unequal datasets. |
Copied to clipboard
| Challenge: | Current methods to improve data quality are labor-intensive or prone to factual errors caused by LLM hallucinations. |
| Approach: | They propose a method which reformats the responses of instruction data into a format that better aligns with pre-established criteria and the collated evidence. |
| Outcome: | The proposed approach minimizes human annotation, hallucination, and the difficulty in scaling, remaining orthogonal to existing alignment techniques. |
Copied to clipboard
| Challenge: | Existing methods use contrastive learning (CL) to learn effective sentence representations, but require extensive human annotation. |
| Approach: | They propose a reinforcement learning approach for fine-tuning small-parameter LLMs to generate high-quality hard contrastive data without human feedback. |
| Outcome: | The proposed method achieves state-of-the-art on seven semantic text similarity tasks. |
Copied to clipboard
| Challenge: | a multi-lingual approach to training dialog systems is expensive and tedious, but it can be useful for cross-lingual support. |
| Approach: | They propose to annotate data for multiple languages and train a multi-lingual dialog system for each language. |
| Outcome: | The proposed framework bypasses the expensive human annotation and achieves promising results. |
Copied to clipboard
| Challenge: | To properly infer the intention of the narrator, one needs a certain degree of common sense and social intuition. |
| Approach: | They propose a task that uses common sense to extract pairs of questions that are appropriate candidates for the task. |
| Outcome: | The proposed method exploits commonalities in experiences people share online to extract pairs of semantically plausible advice-seeking questions that are appropriate candidates for the cloze task. |
Copied to clipboard
| Challenge: | AnaScore metric aims to evaluate the strength of semantic parallelism in sentence analogies. |
| Approach: | They propose an automatic metric to evaluate the strength of semantic parallelism in sentence analogies. |
| Outcome: | The proposed metric shows that formally explainable examples are more beneficial for analogical reasoning, whereas ambiguous analogies with no clear criterion tend to hinder inference. |
Copied to clipboard
| Challenge: | Large Language Models excel at a low-resource level given limited data, but are unsuitable for runtime systems which require low latency. |
| Approach: | They propose a method to augment training data for a model 40x smaller (500M parameters) they use Alexa to generate synthetic data from Alexa 20B to augment the training set . |
| Outcome: | The proposed method improves low-resource SP on two datasets in low-source settings. |
Copied to clipboard
| Challenge: | Existing studies on the ability of a model to make consistently correct predictions in the presence of perturbations have not been conducted in open-domain question answering (OpenQA). |
| Approach: | They propose a query-side contrastive loss to improve the dense passage retriever (DPR) to improve DPR training. |
| Outcome: | The proposed approach improves the density of the dense passage retriever (DPR) training set without sacrificing accuracy on standard test sets. |
Copied to clipboard
| Challenge: | Existing methods for weakly supervised multi-hop pretraining require costly human annotation. |
| Approach: | They propose a method for weakly supervised multi-hop retriever pretraining without human efforts by generating vector representations of complex questions and subquestion as weak supervision for pre-training. |
| Outcome: | The proposed method is effective and robust on limited data and computational resources. |
Copied to clipboard
| Challenge: | Synthetic data generation is an increasingly popular way of training models without the need for large, manually labeled datasets. |
| Approach: | They propose a framework that aligns open-source small models to efficiently generate large-scale embedding data. |
| Outcome: | The proposed framework outperforms state-of-the-art embedding models by using only 1/10 of the GPT API calls. |
Copied to clipboard
| Challenge: | Existing methods to eliminate hallucinations require expensive human annotation . hallucination in multimodal large language models poses unique challenges for current research . |
| Approach: | They propose a fine-grained unlearning framework that performs gradient ascent to eliminate hallucinations without paired data. |
| Outcome: | The proposed method reduces hallucinations while preserving quality with modest computational overhead. |
Copied to clipboard
| Challenge: | Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. |
| Approach: | They hypothesize that context plays an important role in accurate human annotation and add uncertainty measures can improve model accuracy and calibration. |
| Outcome: | The proposed model can be better calibrated by adding uncertainty measures to models with better accuracy and calibration. |
Copied to clipboard
| Challenge: | Figures of speech are ubiquitous in many forms of discourse, allowing people to convey complex, abstract ideas and evoke emotion. |
| Approach: | They develop a dataset for multimodal figurative language understanding using human annotation and an automatic pipeline to generate a multimodal dataset. |
| Outcome: | The proposed dataset performs better than human vision and language models compared with a human dataset . |
Copied to clipboard
| Challenge: | Existing techniques for parsing natural-language utterances are vulnerable to adversarial attacks, requiring large amounts of labelled data and expensive human annotation. |
| Approach: | They propose to enhance the adversarial robustness of a prompt-based semantic parser based on a language model trained on code by constructing a set of demonstration examples. |
| Outcome: | The proposed method can be enhanced without significant amounts of labelled data or expensive human annotations on in-domain semantic parsing data. |
Copied to clipboard
| Challenge: | utilizing human annotations can enhance critique ability, but model-generated critiques suffer from inherent flaws due to complexity of critique . a new framework that leverages multi-agent feedback improves critique ability . |
| Approach: | They propose a framework that leverages multi-agent feedback to improve critique ability . they propose to use supervised fine-tuning and reinforcement learning to improve this capability . |
| Outcome: | The proposed framework improves critique ability in both supervised fine-tuning and reinforcement learning stages. |
Copied to clipboard
| Challenge: | Efforts to improve instruction tuning often focus on higher-quality supervised fine-tuning datasets, typically requiring data filtering with proprietary LLMs or human annotation. |
| Approach: | They propose a Mixup-based recipe that elevates LLM instruction tuning without relying on well-curated datasets. |
| Outcome: | The proposed model improves instruction-following and healthcare-specific tasks with consistent improvements across LLM families and SFT datasets. |
Copied to clipboard
| Challenge: | Variation in human annotation and human perspectives has drawn increasing attention in natural language processing research. |
| Approach: | They propose to use annotation formats that better capture granularity and uncertainty of individual judgments and annotation modeling that leverages socio-demographic features to better represent and predict underrepresented or minority perspectives. |
| Outcome: | The proposed tasks aim to advance natural language processing research towards more faithfully reflecting the diversity of human interpretation, enhancing both inclusiveness and fairness in language technologies. |
Copied to clipboard
| Challenge: | a new commonsense knowledge model, NovaCOMET, combines knowledge and general task models. |
| Approach: | They propose an open commonsense knowledge model that combines knowledge and general task models. |
| Outcome: | The proposed model matches or exceeds existing knowledge models on commonsense reasoning tasks. |
Copied to clipboard
| Challenge: | Sensor names are alphanumeric strings that encode key contextual information such as their function or physical location. |
| Approach: | They propose a self-supervised framework that can learn to segment sensor names without human annotation. |
| Outcome: | The proposed framework can learn to segment sensor names without human annotation on buildings. |
Copied to clipboard
| Challenge: | Existing studies have focused on extracting emotion causes from news articles, but lack of fine-grained annotations has limited the ECE task. |
| Approach: | They propose a new ECE framework that extracts emotion causes from social media data without relying on human annotations. |
| Outcome: | The proposed framework achieves high extraction performance and generalizability without relying on human annotations. |
Copied to clipboard
| Challenge: | Current state-of-the-art methods require expensive human annotation and struggle with domain transfer, limiting their practical deployment. |
| Approach: | They propose a benchmark spanning seven diverse domains to evaluate ATE performance . they propose psuedo-labels and post-hoc heuristics to ensure generalizability . |
| Outcome: | The proposed model outperforms supervised cross-domain encoder models and few-shot learning baselines on the document- and corpus-levels and its GPT-4o teacher on the benchmark. |
Copied to clipboard
| Challenge: | Detecting and identifying events is an important subtask of event extraction. |
| Approach: | They build a large event-related candidate set with good coverage and apply an adversarial training mechanism to iteratively identify informative instances from the candidate set and filter out those noisy ones. |
| Outcome: | The proposed method significantly outperforms the state-of-the-art methods on two real-world datasets. |
Copied to clipboard
| Challenge: | a growing number of generative AI systems are detecting text generated by a model or written by . humans perform poorly at the detection task, but show no significant biases on the studied attributes. |
| Approach: | They examine gender, race/ethnicity, English-language learner status, and economic status . they find several models tend to classify disadvantaged groups as machine-generated . |
| Outcome: | The proposed models show strong performance but can cause negative impacts . the models classify disadvantaged groups as machine-generated, while economically disadvantaged students' essays are less likely to be classified as machine generated . |
Copied to clipboard
| Challenge: | Recent years have positioned Large Language Models (LLMs) as powerful question answering (QA) tools, shifting users away from interacting in communities towards discourse with AI-driven conversational interfaces. |
| Approach: | They propose to use a QA preference dataset to fine-tune and align Large Language Models (LLMs) from more than 7.4 million submissions and 82 million comments from 2008 to 2022 in Reddit’s 15 largest finance communities. |
| Outcome: | The proposed framework improves on the social quality of the data, and the proposed framework is more accurate and more specific. |
Copied to clipboard
| Challenge: | Existing metrics that rely on comparisons to a set of known correct responses do not account for the variety of responses and therefore correlate poorly with human judgment. |
| Approach: | They propose a method of manipulating a golden response to create a new negative response that is designed to be inappropriate within the context while maintaining high similarity with the original golden response. |
| Outcome: | The proposed model can be made using unsupervised learning for the next-utterance prediction task on English datasets and shows that using the negative samples alongside random negative samples can increase the model’s correlation with human evaluations. |
Copied to clipboard
| Challenge: | Current evaluation of German automatic text simplification relies on general-purpose metrics such as SARI, BLEU, and BERTScore. |
| Approach: | They propose a German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency. |
| Outcome: | The proposed metric achieves higher correlations with human judgments than widely used ATS metrics. |
Copied to clipboard
| Challenge: | Existing LLMs lack high-quality data sources and lack robust data filtration strategies. |
| Approach: | They develop a framework to enhance the capabilities of LLM-based agents under data scarcity. |
| Outcome: | The proposed framework improves the capabilities of LLM-based agents under data scarcity. |
Copied to clipboard
| Challenge: | Existing work focuses on learning deep NER models with weak supervision without any human annotation. |
| Approach: | They propose a framework that can suppress the noise of the weak labels and fine-tune over the strongly labeled data. |
| Outcome: | The proposed framework outperforms existing methods on Named Entity Recognition tasks with weak supervision and weakly labeled data. |
Copied to clipboard
| Challenge: | Existing research on information-seeking conversations is stymied by the lack of training data. |
| Approach: | They propose to use autoconv for synthetic conversation generation to capture the characteristics of the information-seeking process and fine tune an LLM with a few human conversations to generate synthetic conversations with high quality. |
| Outcome: | The proposed model improves on two commonly-used datasets and alleviates the dependence on human annotation. |
Copied to clipboard
| Challenge: | Existing data synthesis methods focus on general-purpose tasks and fail to capture domain-specific terminology and reasoning patterns. |
| Approach: | They propose a framework that generates domain-specific instruction datasets without human supervision by pairing task-informed keywords with different cognitive levels from Bloom’s Taxonomy. |
| Outcome: | The proposed framework generates domain-specific instruction datasets without human supervision and achieves significant improvements over existing methods. |
Copied to clipboard
| Challenge: | Existing approaches to zero-shot cross-lingual spoken language understanding rely on shared parameters, which can only perform implicit alignment across languages. |
| Approach: | They propose a global-local contrastive learning framework to achieve a fine-grained cross-lingual transfer . they employ bilingual dictionaries to construct multilingual views of the same utterance . |
| Outcome: | Experiments on MultiATIS++ show that GL-CLeF achieves the best performance . GL is based on dictionaries and encourages representations to be more similar than negative example pairs . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for paraphrase generation are not designed for the task, but adopted from other evaluation tasks. |
| Approach: | They propose a new evaluation metric for paraphrase generation that uses reference-based and reference-free metrics. |
| Outcome: | The proposed evaluation metric outperforms existing metrics and is more reliable than reference-based metrics. |
Copied to clipboard
| Challenge: | Existing studies have used the correlation information stored in samples for self-supervised learning, but they feed the training pairs in a random order without consideration of difficulty. |
| Approach: | They propose to inject curriculum learning into weakly supervised multimodal correlation learning by scoring and feeding pairs according to difficulty. |
| Outcome: | The proposed model achieves state-of-the-art on multimodal sentiment analysis without human annotation. |
Copied to clipboard
| Challenge: | Recent work on textual Aspect-Based Sentiment Analysis (ABSA) has demonstrated promising performance, but limited semantics derived from raw data. |
| Approach: | They propose a method that provides visual semantics to reinforce textual ABSA by adding additional augmentations to the input data. |
| Outcome: | The proposed method can provide visual semantics to reinforce the textual extraction. |
Copied to clipboard
| Challenge: | Current approaches to medical entity retrieval generalize poorly to unseen sub-specialties . zero-shot retrieval is challenging due to the high degree of ambiguity and variability in medical corpora . |
| Approach: | They propose a set of learning tasks designed to train efficient zero-shot entity retrieval models. |
| Outcome: | The proposed architecture outperforms common zero-shot benchmarks with 7% to 30% higher recall across multiple major medical ontologies. |
Copied to clipboard
| Challenge: | Existing evaluations assess static recall or isolated visual grounding, leaving unanswered whether VLMs possess robust and transferable cultural understanding. |
| Approach: | They propose a multimodal, multicultural benchmark to evaluate the robustness of everyday cultural knowledge in vision-language models across linguistic rephrasings and visual modalities. |
| Outcome: | ‘BLEnD-Vis‘ constructs 313 culturally grounded question templates spanning 16 regions and generates three aligned multiple-choice formats. |
Copied to clipboard
| Challenge: | Existing methods to extract relation facts from limited labeled corpora are laborintensive to obtain . Existing approaches use self-training to generate pseudo labels that will cause gradual drift problem or leverage meta-learning scheme which does not solicit feedback explicitly. |
| Approach: | They propose a Gradient Imitation Reinforcement Learning method to encourage pseudo label data to imitate gradient descent direction on labeled data and bootstrap its optimization capability through trial and error. |
| Outcome: | The proposed method handles two major scenarios in low-resource relation extraction when no unlabeled data is available. |
Copied to clipboard
| Challenge: | In this paper, we aim to generate text classification data given arbitrary class definitions . Traditional supervised text classification fine-tunes models on expensive human annotation . |
| Approach: | They propose a framework that can generate text classification data given arbitrary class definitions . they use instruction-to-data mappings and in-context augmentation to refine the framework . |
| Outcome: | The proposed framework outperforms existing methods on benchmarks and training data generation by prompt engineering. |
Copied to clipboard
| Challenge: | Existing methods for OCR correction are mostly supervised methods that correct recognition errors in a single output. |
| Approach: | They propose a sequence-to-sequence model with attention and a decoder with attention averaging to search for consensus among multiple sequences. |
| Outcome: | The proposed methods cut the character and word error rates nearly in half on single inputs and can rival supervised methods. |
Copied to clipboard
| Challenge: | Prior work has found that language models (LMs) can harm users in hard-to-predict ways, and human annotation is expensive, limiting the number and diversity of test cases. |
| Approach: | They propose to generate test inputs using an LM itself, and use a classifier to detect harmful behavior on test input. |
| Outcome: | The proposed approach detects tens of thousands of offensive responses in a 280B parameter LM chatbot. |
Copied to clipboard
| Challenge: | Document-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document. |
| Approach: | They propose to transfer an English document to Japanese to promote DocRE in other languages. |
| Outcome: | The proposed model reduces the human edit steps by 50% compared with the previous approach. |
Copied to clipboard
| Challenge: | Recent research has focused on identifying text that introduces new, previously unknown information, but has seen a decline in novelty detection due to the rise of large language models. |
| Approach: | They propose a novel automated metric for evaluating document-level novelty that aggregates the novelty and salience scores of atomic information and provides high interpretability and a detailed analysis of a document's novelty. |
| Outcome: | The proposed metric scores high on the TAP-DLND 1.0 dataset and a human-annotated dataset. |
Copied to clipboard
| Challenge: | a recent study shows that reward models overfit on superficial features, hindering generalization performance . prevailing approach to training preference-based reward models presents several challenges . |
| Approach: | They propose a method that uses synthetic natural language critiques to provide additional feedback to large language models. |
| Outcome: | The proposed approach improves performance and data efficiency of RMs initialized from different pretrained models, reducing the reliance on costly human annotations. |
Copied to clipboard
| Challenge: | Existing methods to generate large-scale datasets are difficult in closed domains where human annotation requires domain expertise. |
| Approach: | They propose a method to generate diverse and semantic questions in a low-resource setting with the aim of summarizing healthcare questions. |
| Outcome: | The proposed method generates diverse, fluent, and informative summarized questions on healthcare question summarization datasets. |
Copied to clipboard
| Challenge: | Argumentation plays a central role in human communication, where refuting or attacking others’ arguments is a common persuasion strategy. |
| Approach: | They propose a novel annotation scheme that captures common modes and complex rhetorical moves in attacks along with the implicit presuppositions and value judgments. |
| Outcome: | The proposed scheme shows moderate agreement between the two annotations, indicating that human annotation is feasible. |
Copied to clipboard
| Challenge: | Recent advances in Vision-Language Models and the scarcity of high-quality multi-modal alignment data have inspired numerous researches on synthetic VLM data generation. |
| Approach: | They propose a multi-modal data construction pipeline that organizes the final output into a Python code format. |
| Outcome: | The proposed pipeline improves visual question answering and visual grounding benchmarks across different VLMs. |
Copied to clipboard
| Challenge: | Existing models for dialogue empathy focus on the emotion flow in one direction, from context to response. |
| Approach: | They propose a dual-generative model to construct emotional consensus and use unpaired data to produce pseudo paired empathetic samples. |
| Outcome: | The proposed model outperforms baseline models in producing coherent and empathetic responses. |
Copied to clipboard
| Challenge: | a dataset for scientific entity extraction, classification, and resolution has been developed . a generic conceptual formalism for scientific entities is feasible, the authors say . |
| Approach: | They propose a STEM-ECR dataset that provides a domain-independent benchmark for scientific entity extraction, classification, and resolution tasks. |
| Outcome: | The proposed dataset provides a benchmark for evaluation of scientific entity extraction, classification, and resolution tasks in a domain-independent fashion. |
Copied to clipboard
| Challenge: | Active learning (AL) is a human-and-model-in-the-loop paradigm that iteratively selects informative unlabeled data for human annotation. |
| Approach: | They propose to simulate active learning by using an already labeled dataset as the pool of unlabeled data. |
| Outcome: | The proposed model-in-the-loop paradigm can be used to perform experiments with human annotations on-the fly. |
Copied to clipboard
| Challenge: | Existing fact-checking systems struggle with attribution quality, as their generated explanations can include hallucinations. |
| Approach: | They propose a protocol to assess attribution quality in fact-checking explanations using human annotation and automatic annotation. |
| Outcome: | The proposed protocol can be automated, the authors show . best-performing LLMs still generate explanations that are not always accurate . |
Copied to clipboard
| Challenge: | Existing methods for creating a vision question-answering with natural language explanations rely on human annotations that are time-consuming and costly. |
| Approach: | They propose a method that generates high-quality natural language explanations using LVLMs by using visual prompts. |
| Outcome: | The proposed method generates high-quality synthetic VQA-NLE datasets 20x faster than human annotations with minimal decrease in qualitative metrics. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in answering questions pertaining to commonsense reasoning and inference. |
| Approach: | They prompt LLMs to generate items in the style of a benchmark for commonsense reasoning . they find that LLM authors that answer COPA items are more successful . |
| Outcome: | The authors' responses to their own items and their own generated items are better than those of the original LLMs. |
Copied to clipboard
| Challenge: | Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user’s input text. |
| Approach: | They propose a training algorithm that trains T2I models to be faithful to the input text. |
| Outcome: | The proposed model improves both the semantic alignment and aesthetic appeal of two diffusion-based T2I models, evidenced by multiple benchmarks (+1.7% on TIFA, +2.9% on DSG1K, +3.4% on VILA aesthetic). |
Copied to clipboard
| Challenge: | Norwegian is under-represented within the most impressive breakthroughs in NLP tasks. |
| Approach: | they investigate the impact of existing Norwegian language models on Norwegian generation tasks . they pre-trained 4 Norwegian Open Language Models from parameter scales and architectures . |
| Outcome: | The proposed benchmark evaluates the performance of language models on Norwegian generation tasks. |
Copied to clipboard
| Challenge: | Existing methods for word embedding evaluation are computationally expensive and task-specific. |
| Approach: | They propose a minimally supervised method for generating word embedding evaluation datasets for a large number of languages using existing dependency treebanks and parsers. |
| Outcome: | The proposed method evaluates three popular word embedding algorithms against these datasets and shows that their performance varies between syntactic categories. |
Copied to clipboard
| Challenge: | Social media advertising allows entities to construct narratives that align with their commercial interests and sway public perception. |
| Approach: | They propose to classify climate-related narratives into seven categories based on existing definitions and data. |
| Outcome: | The proposed method outperforms other methods and can reduce human annotation costs. |
Copied to clipboard
| Challenge: | Semi-supervised learning (SSL) is a popular technique for reducing the reliance on human annotations for NLI tasks. |
| Approach: | They propose a way to incorporate unlabeled data into semi-supervised learning (SSL) using a conditional language model, they propose to generate hypotheses for unlabed sentences . |
| Outcome: | The proposed framework significantly improves the performance of four NLI datasets in low-resource settings. |
Copied to clipboard
| Challenge: | Generalized Category Discovery (GCD) is a practical and challenging open-world task that aims to recognize both known and novel categories in unlabeled data using limited labeled data from known categories. |
| Approach: | They propose a framework for generalized category discovery that actively learns from diverse and collaborative feedback. |
| Outcome: | The proposed framework improves instance-level contrastive features, generates category descriptions, and aligns uncertain instances with LLM-selected category descriptions. |
Copied to clipboard
| Challenge: | Using the AnnCor CHILDES Treebank, we assign adult grammar syntactic structures to children's utterances. |
| Approach: | They propose a partially manually verified treebank for Dutch CHILDES corpora . they argue that human annotation and automatic checks on this annotation must go hand in hand . |
| Outcome: | The AnnCor CHILDES Treebank is the first partially manually verified treebank for Dutch CHILdes corpora. |
Copied to clipboard
| Challenge: | Existing datasets for detecting online propaganda use weak labels that can be noisy and incorrect. |
| Approach: | They propose a dataset for detecting online propaganda with high-quality labels . they show that state-of-the-art language models fail in detecting propaganda when trained with weak labels compared to prompt-based learning . |
| Outcome: | The proposed dataset is the first large-scale dataset for detecting online propaganda that was created through human annotation. |
Copied to clipboard
| Challenge: | Existing CRS datasets focus on immediate requests from users, while lack proactive guidance to the recommendation scenario. |
| Approach: | They propose a topic-guided conversational recommendation dataset . it incorporates topic threads to enforce natural semantic transitions towards the recommendation scenario . |
| Outcome: | The proposed approach is more reasonable and controllable than previous approaches. |
Copied to clipboard
| Challenge: | Logical table-to-text generation requires models to derive logical-level facts from table records via logical inference. |
| Approach: | They propose a pretrained logical form generator framework to improve generation fidelity . they use a dataset to test the logical inference accuracy of the framework . |
| Outcome: | The proposed framework outperforms baselines on LOGICNLG and CONTLOG on two benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to enlarge SLU data require large amounts of labelled data. |
| Approach: | They propose a data augmentation method with atomic templates for Spoken Language Understanding which generates atomic exemplars from atomic template. |
| Outcome: | The proposed method improves on a DSTC 2&3 dataset which is a domain adaptation setting of SLU. |
Copied to clipboard
| Challenge: | Logic-based approaches to reasoning have lost popularity due to limited scalability and coverage. |
| Approach: | They present a dataset of 28K sentence-level NL-FOL pairs from GPT4 and a LogicLLaMA2-7B/13B fine-tuned on MALLS for NL translation. |
| Outcome: | The proposed model can be used standalone or to correct previously generated rules by GPT3.5. |
Copied to clipboard
| Challenge: | Recent studies have raised concerns about the potential threats large language models pose to academic integrity and copyright protection. |
| Approach: | They propose a dataset of 46.5K synthetic text pairs that represent three major types of plagiarism: verbatim copying, paraphrasing, and summarization. |
| Outcome: | The proposed dataset shows that GPT-3.5 Turbo can produce high-quality paraphrases and summaries without significantly increasing text complexity compared to GPT-4 Turbo. |
Copied to clipboard
| Challenge: | Existing approaches for constructing PRM training data rely on human annotation or sampling-based labeling methods that require repeated LLM calls. |
| Approach: | They propose a framework that synthesizes PRM training data by annotating step-level error labels using formal verification tools such as Z3 and Isabelle. |
| Outcome: | The proposed framework synthesizes PRM training data from formal logic and theorem proving tasks without human annotation or additional LLM calls. |
Copied to clipboard
| Challenge: | Existing methods for annotating instruction data are expensive and difficult to scale. |
| Approach: | They propose a method to automatically build instruction data from an unlabeled corpus without heavy reliance on proprietary LLMs and human annotation. |
| Outcome: | The proposed method outperforms existing methods on AlpacaEval leaderboard and other open-source methods. |
Copied to clipboard
| Challenge: | Prior work on instruction tuning relies on expensive human annotation and crowd-sourced datasets with alignment issues. |
| Approach: | They propose a method to generate instructions via LLMs from human-written corpus examples using reverse instructions. |
| Outcome: | The proposed method outperforms larger language models without instruction tuning on tasks such as story/recipe generation and long-form question answering. |
Copied to clipboard
| Challenge: | Existing approaches for self-supervision operate at word form level, which serves as a surrogate for the underlying semantic content. |
| Approach: | They propose a method to employ weak-supervision directly at the word sense level, without the use of human annotation. |
| Outcome: | The proposed model achieves significantly improved lexical understanding without human annotation on the ‘Word in Context’ task. |
Copied to clipboard
| Challenge: | Existing sentiment lexicons are compiled by (machine) translation from English resources, obscuring language-specific characteristics of sentiment-loaded vocabulary. |
| Approach: | They propose a gold standard for sentiment annotation of Swedish terms using the SALDO lexicon and the Gigaword corpus. |
| Outcome: | The proposed model is based on the free SALDO lexicon and the Gigaword corpus and is compared with existing models using human annotations. |
Copied to clipboard
| Challenge: | a method for process supervision has shown significant improvements in multi-step problem solving . despite the advances in process supervision, there are still easily observable mistakes in state-of-the-art LLMs. |
| Approach: | They propose a method for automating data curation by using a trained verifier to evaluate intermediate steps generated by a reasoner. |
| Outcome: | The proposed method improves the performance of PaLM 2 on math and coding tasks. |
Copied to clipboard
| Challenge: | Existing methods for acquiring large-scale intentions generate product-centric intentions without product images and incur high costs for scalability. |
| Approach: | They propose a multimodal framework that allows Large Vision-Language Models to infer purchase intentions from multimodal product metadata and prioritize human-centric ones. |
| Outcome: | The proposed framework shows that it is robust to different prompts and superior to previous methods. |
Copied to clipboard
| Challenge: | Existing knowledge graphs lack two desired features for modeling entity relationships: openness and informativeness. |
| Approach: | They propose a self-supervised learning method to extract relation descriptions with the analysis of dependency patterns and generate relation descriptions using a transformer-based relation description synthesizing model. |
| Outcome: | The proposed system extracts and generates high-quality relation descriptions without human labeling. |
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs struggle with instructions containing multiple constraints. |
| Approach: | They propose a self-correction pipeline that decomposes the original instruction into a list of constraints and uses a Critic model to decide when and where the LLM’s response needs refinement. |
| Outcome: | The proposed model outperforms GPT-4 on RealInstruct and IFEval even with weak feedback. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on complex methodologies like Proximal Policy Optimization (PPO) that require extensive hyper-parameter tuning and pose challenges in sample efficiency and stability. |
| Approach: | They propose an innovative framework that leverages direct preference optimization techniques but extends them by estimating the conditionally optimal policy directly from the model’s responses. |
| Outcome: | The proposed framework matches and exceeds the effectiveness of Proximal Policy Optimization (PPO) in terms of convergence speed and alignment of model responses with human preferences. |
Copied to clipboard
| Challenge: | Existing learning metrics are limited to tasks where large human ratings are available. |
| Approach: | They propose a model-based natural language generation (NLG) evaluation metric that is highly correlated with human judgements without requiring human annotation. |
| Outcome: | The proposed metric outperforms all prior unsupervised metrics on multiple NLG tasks including translation, image captioning, and WebNLG text generation. |
Copied to clipboard
| Challenge: | Existing evaluation methods rely on human judgment to assess data accuracy and visual communication, which is costly and unscalable. |
| Approach: | They propose a framework that leverages Visual Question Answering (VQA) models to automate the evaluation of LLM-generated data visualizations. |
| Outcome: | The proposed framework assesses data representation quality and communicative clarity of charts using two leading VQA benchmark datasets, ChartQA and PlotQA, with visualizations generated by OpenAI’s GPT-3.5 Turbo and Meta’s Llama 3.1 70B-Instruct models. |
Copied to clipboard
| Challenge: | Recent progress in large language models is driven by scaling of training compute through pre-training with nexttoken prediction (NTP) or post-training (RL) Pre-training using NTP enables models to acquire extensive knowledge and skills from general data, but it suffers from data inefficiency and catastrophic forgetting in continual learning settings. |
| Approach: | They propose to scale training compute through pre-training with next-token prediction (NTP) or post-training by scaling reinforcement learning (RL) to improve learning from general data. |
| Outcome: | Experiments on multiple benchmarks and models show that the proposed approach improves continual pre-training and provides a strong foundation for post-training on Qwen3-8B-Base. |
Copied to clipboard
| Challenge: | Named entity recognition models rely on large amounts of labeled data, making them challenging to extend to new, lower-resource languages. |
| Approach: | They propose a method for bootstrapping named entity recognition models in under-resourced languages . they use cross-lingual transfer learning and targeted annotation of only uncertain entities . |
| Outcome: | The proposed method achieves competitive accuracy with just one-tenth of training data. |
Copied to clipboard
| Challenge: | Existing approaches to learn dialogue discourse parsing with related tasks require additional annotation, thus limiting their generality. |
| Approach: | They propose a multitasking framework that integrates dialogue discourse parsing with addressee recognition to reflect relation-based structure of dialogue. |
| Outcome: | The proposed framework outperforms baselines on the Molweni and STAC datasets. |
Copied to clipboard
| Challenge: | Recent advances in language generation models can be used to assist users in a variety of tasks, but there are risks associated with introducing LLM biases into consequential decisions. |
| Approach: | They propose to use a template-generated dataset to measure subtler correlated decisions that LLMs make between social groups and unrelated positive and negative attributes. |
| Outcome: | The proposed model can be used to evaluate progress in more generalized biases and extend the benchmark with minimal human annotation. |
Copied to clipboard
| Challenge: | Existing LLMs lack datasets and biased training tasks to follow speech instructions. |
| Approach: | They propose a query rewriting framework that uses multiple agents to annotate and validate the synthesized speech. |
| Outcome: | The proposed framework can transform text instructions into distributions more suitable for TTS models for speech synthesis without human annotation. |
Copied to clipboard
| Challenge: | Acceptability is one of the general language understanding evaluation benchmarks (GLUE) probing tasks . EsCoLA consists of 11,174 sentences and their acceptability judgements as found in well-known Spanish reference grammars. |
| Approach: | They propose to use a corpus of linguistic acceptability (ESCoLA) EsCoLA consists of 11,174 sentences and their acceptability judgements . |
| Outcome: | The proposed task is based on 11,174 sentences and their acceptability judgements as found in well-known Spanish reference grammars. |
Copied to clipboard
| Challenge: | Existing studies on relation extraction (RE) use labeled training data for relation extraction models but it is expensive and time-consuming. |
| Approach: | They propose a dual supervision framework which utilizes both types of data to train relation extraction models. |
| Outcome: | The proposed framework can predict labels by human annotation and distant supervision without labeling bias since it is expensive and time-consuming. |
Copied to clipboard
| Challenge: | Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. |
| Approach: | They propose a framework that identifies and intervenes on the specific attention heads causally responsible for toxic generation. |
| Outcome: | The proposed framework reduces toxic generation by 5.34% while preserving linguistic fluency and speeding up head selection. |
Copied to clipboard
| Challenge: | Existing methods for measuring identity fusion are limited and require controlled surveys or direct field contact. |
| Approach: | They propose a new metric that integrates cognitive linguistics with large language models to measure identity fusion. |
| Outcome: | The proposed metric outperforms existing methods and human annotations in violence risk assessment. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have greatly expanded the scope of legal AI. |
| Approach: | They propose a method that generates questionnaires to help users refine queries . they leverage an iterative training process that collects valuable questionnaires . |
| Outcome: | The proposed method improves the completeness of queries and ensures the performance of domain-specific models in downstream legal tasks. |
Copied to clipboard
| Challenge: | Existing methods for enhancing dialogue performance rely on summarizing behavior . e-commerce chatbots need to align their dialogue strategies with human behavior to achieve coherent, human-like conversations with customers. |
| Approach: | They propose a method to extract core patterns from dialogue data and integrate them into models by mining service thought processes using a multi-agent aPproach. |
| Outcome: | The proposed method outperforms manual methods and outperfies baselines on Taobao in China. |
Copied to clipboard
| Challenge: | Understanding users’ contextual search intent when generating responses is an understudied topic for conversational question answering (QA). |
| Approach: | They propose a method that allows LLMs to decide when to retrieve in RAG settings given a conversational context. |
| Outcome: | The proposed method improves on three conversational QA datasets and criticizes the quality of generated responses. |
Copied to clipboard
| Challenge: | Existing annotations for irony are difficult, and the recognition of it is difficult due to its polarity. |
| Approach: | They propose a fine-grained annotation scheme centered on irony that highlights the tokens responsible for its activation and their morpho-syntactic features. |
| Outcome: | The proposed scheme highlights the tokens responsible for irony activation and their morpho-syntactic features. |
Copied to clipboard
| Challenge: | a new study examines the performance of code-switching IR in monolingual contexts . code-witching is a pervasive linguistic phenomenon in global communication . |
| Approach: | They propose a benchmark to evaluate code-switching IR in monolingual contexts . they propose CS-MTEB, which measures performance declines of up to 27% . |
| Outcome: | The proposed benchmark shows that code-switching performance is degraded by 27% . the proposed benchmark is based on a dataset of mixed-language queries . |
Copied to clipboard
| Challenge: | Existing approaches to augment Large Language Models (LLMs) with computational capabilities have focused on short Chain-of-thought (CoT) integrating tool-use into long CoT remains underexplored due to the scarcity of training data and the challenge of integrating it without compromising the model’s intrinsic long-chain reasoning. |
| Approach: | They propose a framework that enables spontaneous tool-use during long CoT reasoning without additional human annotation. |
| Outcome: | Experiments on AIME and GPQA-Diamond show that DART significantly outperforms existing methods, successfully harmonizing tool execution with long CoT reasoning. |
Copied to clipboard
| Challenge: | Using a multi-layered scheme for the fine-grained annotation of irony on Italian Twitter is a challenging task to be performed by both human annotators and automatic NLP systems. |
| Approach: | They propose to apply a multi-layered scheme for the fine-grained annotation of irony to an Italian Twitter corpus. |
| Outcome: | The proposed scheme can be validated on Italian irony-laden social media contents and is available in the cross- and multi-lingual perspective. |
Copied to clipboard
| Challenge: | Evaluating conversational information retrieval systems requires a significant amount of human labor for annotation. |
| Approach: | They propose to use human annotation to calibrate evaluation results to eliminate evaluation biases. |
| Outcome: | The proposed method consumes less than 1% of human labor and achieves a consistency rate of 95%-99% with human evaluation results. |
Copied to clipboard
| Challenge: | Existing knowledge resources for sentiment analysis (SA) tasks are either large, common-sense knowledge graphs (KGs) that cover a limited amount of polarities/emotions or they are smaller in size (e.g. lexicons) . however, these resources are limited by the low coverage of e.t. and scalability. |
| Approach: | They propose a new directed KG called ‘RELATE’ which incorporates the benefit of semantics without relying on costly human annotation. |
| Outcome: | The proposed KG overcomes low coverage of emotions and scalability issues . it is the first KG of its size to cover Ekman’s six basic emotions that are directed towards entities. |
Copied to clipboard
| Challenge: | Entity linking models have been successful in capturing semantic features, but the NIL prediction problem has not been addressed. |
| Approach: | They propose an entity linking dataset that categorizes mentions linking to NIL into Missing Entity and Non-Entity Phrases. |
| Outcome: | The proposed dataset categorizes mentions linking to NIL into Missing Entity and Non-Entity Phrase categories and ensures the presence of mentions by human annotation and entity masking. |
Copied to clipboard
| Challenge: | Distant supervision reduces the reliance on human annotation in named entity recognition tasks. |
| Approach: | They propose a class-rebalancing self-training framework for improving distantly-supervised named entity recognition by using a flexible threshold and a hybrid pseudo label. |
| Outcome: | The proposed model achieves state-of-the-art on five flat and two nested datasets and compares with other methods on the same dataset. |
Copied to clipboard
| Challenge: | Evaluating open-domain dialogue systems is challenging because of the one-to-many problem. |
| Approach: | They propose a reference-based dialogue evaluation approach that leverages the pre-created utterance as reference other than the gold response to relieve the one-to-many problem. |
| Outcome: | The proposed method outperforms state-of-the-art evaluation methods on three datasets and two existing benchmarks. |
Copied to clipboard
| Challenge: | Existing knowledge retrieval methods for task-oriented dialogues are limited by data scarcity and lack of data to annotate. |
| Approach: | They propose an LLM-enhanced model of query-guided knowledge retrieval for task-oriented dialogue . they propose to select the most relevant knowledge from retrieved top-K records and incorporate them as prompts to guide a generator in response generation. |
| Outcome: | The proposed model outperforms state-of-the-art in three benchmarks on three standard benchmarks. |
Copied to clipboard
| Challenge: | In this paper, we examine the role of conversational context in abusive language detection . prior studies have ignored the contextual nature of abusive language, ignoring this aspect . toxicity, hate speech, harmful stereotypes are among the forms of harmful language . |
| Approach: | They propose to use conversational context to analyze abusive language detection using two methods . they use "abusive language" as an umbrella term to refer to various forms of harmful language . |
| Outcome: | The proposed approach is based on two datasets in English and a new dataset of French tweets annotated for hate speech and stereotypes. |
Copied to clipboard
| Challenge: | Scientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers . previous attempts failed to scale up due to heavy human annotation and domain expertise . |
| Approach: | They propose a method to automatically extract TLDR summaries from scientific papers . they propose 'citeSum' with no human annotation, which is 30 times larger than SciTLDR . |
| Outcome: | The proposed approach outperforms most fully-supervised methods on SciTLDR without fine-tuning and achieves state-of-the-art results with only 128 examples. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on textual data, single-document comprehension, or evaluating retrieval and generation in isolation. |
| Approach: | They propose a multimodal RAG benchmark featuring multi-type queries over visually rich document corpora. |
| Outcome: | The proposed benchmark outperforms existing benchmarks in visual retrieval and human-verified queries. |
Copied to clipboard
| Challenge: | Current benchmarks for social biases have limitations in scope, grounding, quality and human effort required. |
| Approach: | They propose to use a language model to help with the development of bias benchmarks . they extend previous work to a new community and set of biases: the Jewish community and antisemitism . |
| Outcome: | The proposed LLM does not perform well on the Jewish community and antisemitism task. |
Copied to clipboard
| Challenge: | a tool for automatic marking up of quantifiers is proposed for Polish . it is trained on a recently annotated corpus of Polish quantificational expressions . |
| Approach: | They propose to use a BERT based neural model to mark up quantifiers in text . they analyse a manually annotated corpus of Polish quantificational expressions and compare it to a human annotation model. |
| Outcome: | The proposed model can be used to build semantically annotated quantifier corpora for other languages. |
Copied to clipboard
| Challenge: | a recent study found that finetuned language models rely on spurious patterns in training data . this limitation limits their performance on out-of-distribution (OOD) test data. |
| Approach: | They propose a method that only requires annotation of a small fraction of training data . they add 1% manual counterfactuals to training data and generate extra counterfacts in vector space . |
| Outcome: | The proposed approach improves sentiment classification using IMDb data and other sets for OOD tests. |
Copied to clipboard
| Challenge: | Experimental results show that this iterative approach leads to consistent improvements in both the policy model and reward model. |
| Approach: | They propose a method that iteratively improves both the policy model and reward model without requiring additional human annotation. |
| Outcome: | The proposed method improves both the policy model and reward model without human annotation. |
Copied to clipboard
| Challenge: | Existing evaluations of commonsense for large language models focus on downstream knowledge tasks, failing to probe whether LLMs truly understand and utilize knowledge or merely memorize it. |
| Approach: | They propose to automatically construct a large benchmark named CoCo which measures LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks. |
| Outcome: | The proposed benchmark systematically assesses LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks. |
Copied to clipboard
| Challenge: | a number of studies have evaluated user satisfaction estimation in TOD systems . current benchmarks for user satisfaction estimates are highly skewed towards dialogues for which the user is satisfied. |
| Approach: | They leverage large language models to generate satisfaction-aware counterfactual dialogues to augment original dialogues of a test collection. |
| Outcome: | The proposed models show higher robustness to increase in dissatisfaction labels than fine-tuned models. |
Copied to clipboard
| Challenge: | Temporal Logic (TL) can be used to specify complex high-level specifications for systems in many engineering domains. |
| Approach: | They propose a framework for translation between NL and TL using Large Language Models . they use a dataset to create a model with 23K NL-TL pairs and human annotation . |
| Outcome: | The proposed framework achieves higher accuracy (> 95%) using only 10% training data compared with baseline model. |
Copied to clipboard
| Challenge: | Existing benchmarks that focus on knowledge-intensive tasks do not reflect diverse educational scenarios. |
| Approach: | They propose a benchmark that incorporates 9 major scenarios and 4,000 educational contexts. |
| Outcome: | The proposed model performs comparable to state-of-the-art large models on the test set. |
Copied to clipboard
| Challenge: | Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. |
| Approach: | They evaluate three annotation setups to integrate comparative judgment into human annotation for machine translation. |
| Outcome: | The proposed approach improves inter-annotator agreement and stability of the annotations. |
Copied to clipboard
| Challenge: | MT-RewardTree provides a framework for constructing, evaluating, and deploying process reward models in machine translation (MT) |
| Approach: | They propose a method for automatically generating token-level preference pairs using approximate Monte Carlo Tree Search. |
| Outcome: | The proposed framework achieves state-of-the-art performance in token-level evaluation and sequence-level analysis. |
Copied to clipboard
| Challenge: | Using a large language model, idea-buckets are automatically retrieved and a small number of participants are able to score the idea without human annotation. |
| Approach: | They propose a large-scale, psychometrically validated system for frequency-based originality scoring that integrates a Large Language Model with externally orchestrated retrieval. |
| Outcome: | The proposed system matches human annotations in idea clustering structure and participant-level scoring while showing strong convergent and external validity. |
Copied to clipboard
| Challenge: | Existing studies on machine translation evaluation focused on quality of individual sentences, while neglecting the importance of contextual information. |
| Approach: | They propose a context-aware machine translation evaluation metric called Cont-COMET . they use the COMET framework to consider the preceding and subsequent contexts of the sentence . |
| Outcome: | The proposed metric improves system-level and segment-level evaluations on the official WMT framework. |
Copied to clipboard
| Challenge: | Language detoxification involves removing toxicity from offensive language. |
| Approach: | They propose an automated pipeline to generate offensive language with implicit offensiveness and trend-aligned slang. |
| Outcome: | The proposed dataset exhibits high pair consistency and greater implicit offensiveness compared to existing Korean datasets and demonstrates applicability to other languages. |
Copied to clipboard
| Challenge: | Existing captioning models ignore existing alt-text metadata and lack transparency if training data is unknown. |
| Approach: | They propose an approach to edit and re-align alt-texts associated with images using human annotation. |
| Outcome: | The proposed approach improves image captions and improves text-to-image generation and zero-shot image classification tasks. |
Copied to clipboard
| Challenge: | Rhetorical strategies are important to persuasive communication, but their analysis relies on human annotation, which is costly, inconsistent and difficult to scale. |
| Approach: | They propose a framework that leverages large language models to generate and label debate data . they fine-tune transformer-based classifiers on this dataset and validate it against human data a . |
| Outcome: | The proposed model achieves high performance and strong generalization across topical domains. |
Copied to clipboard
| Challenge: | Modern language models rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors, but they remain vulnerable to adversarial attacks due to three key limitations: (1) the inefficiency and high cost of human annotation; (2) the vast diversity of potential adversarials; and (3) the risk of feedback bias and reward hacking. |
| Approach: | They propose an iterative adversarial training method that incorporates three key innovations to address these challenges. |
| Outcome: | Experiments on Mistral-7B-Instruct-v0.3 show that the proposed method significantly enhances robustness and reduces harmful outputs from 5.88% to 0.43%. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to remarkable progress across a wide range of natural language processing tasks. |
| Approach: | They propose a training framework that enables fine-tuning LLM agents without human annotation. |
| Outcome: | The proposed framework enables fine-tuning LLM agents without human annotation. |
Copied to clipboard
| Challenge: | Existing methods for extracting chemical procedures from literature are insufficient and low-quality due to the inherent ambiguity of chemical language and the high cost of human annotation. |
| Approach: | They propose a fully fine-tuned large language model (LLM) as a chemical executor to convert between unstructured experimental procedures and structured action sequences. |
| Outcome: | The proposed model outperforms the baseline model on R2D and D2A tasks by 10%. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have aimed to refine their capacity to accurately follow human instructions and navigate intricate scenarios. |
| Approach: | They propose a method that uses a set of instructions to translate English into Japanese and then generates Japanese instruction data using GPT-4. |
| Outcome: | The proposed method outperforms Japanese-Alpaca models in the evaluation benchmarks without human references. |
Copied to clipboard
| Challenge: | Large language models generate preferred responses while avoiding harmful or inappropriate outputs, despite their ability to generate cross-language transferability. |
| Approach: | They introduce the first Polish preference dataset PLLuM-Align, created entirely through human annotation to reflect Polish language and cultural nuances. |
| Outcome: | The proposed dataset lays the groundwork for more aligned Polish LLMs and contributes to the broader goal of multilingual alignment in underrepresented languages. |
Copied to clipboard
| Challenge: | Existing methods for estimating the cognitive complexity of reading comprehension items are expensive, time-consuming, and subject to rater variability. |
| Approach: | They propose to use two dimensions to estimate cognitive complexity of RC items to focus on evidence Scope and transformation level to estimate the cognitive complexity. |
| Outcome: | The proposed models can estimate the cognitive complexity of items by focusing on two dimensions—Evidence Scope and Transformation Level—that indicate the degree of cognitive burden involved in reasoning about the answer. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can reveal toxic or offensive content inadvertently or intentionally. |
| Approach: | They propose to control the diversity of both sides according to the number of samples for fine-tuning, which can directly reflect their impact. |
| Outcome: | The proposed approach improves the performance of large language models after fine-tuning. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) have advanced from perception tasks to complex multi-step reasoning. |
| Approach: | They propose a framework that integrates reinforcement learning with verifiable rewards with process-level supervision through automatically collected rubric-based generative rewards. |
| Outcome: | The proposed framework achieves state-of-the-art performance on six multimodal reasoning benchmarks and significantly improves reasoning faithfulness in dedicated evaluations. |
Copied to clipboard
| Challenge: | Existing datasets for multiword expressions are inconsistently annotated, limited to a single type of MWE, or limited in size. |
| Approach: | They propose to use a new interface to generate MWE annotations for the first time in a dataset of MWE identification. |
| Outcome: | The proposed model outperforms existing models on the DiMSUM dataset. |
Copied to clipboard
| Challenge: | Existing benchmarks on multi-hop QA focus on single-hop and layered ambiguity, but they focus on ambiguous questions . ambiguities can arise at any stage, complicating the reasoning process . |
| Approach: | They propose a benchmark to evaluate ambiguity in multi-hop question answering . they propose MARCH, which uses 2,209 carefully annotated questions . |
| Outcome: | The proposed framework outperforms existing approaches and significantly outperfies existing frameworks. |
Copied to clipboard
| Challenge: | Existing writing assistants rely on supervised fine-tuning to optimize models for multiple revisions. |
| Approach: | They propose a framework that enhances WA performance with rationale and alignment. |
| Outcome: | The proposed framework outperforms state-of-the-art WAs and the closed-source GPT-4o by 3.9 and 7.1 points on average across eight well-established writing-related test sets. |
Copied to clipboard
| Challenge: | Existing approaches of aligning large language models to follow user instructions can lead to undue emphasis on irrelevant documents, which in turn reduces the quality of responses. |
| Approach: | They propose to use a framework to automatically generate high-quality attributed query-response pairs for both supervised fine-tuning and preference optimization stages without human annotation. |
| Outcome: | The proposed framework can generate high-quality attributed query-response pairs without human annotation without human intervention. |
Copied to clipboard
| Challenge: | Recent methods address Chinese Spelling Correction (CSC) with either BERT-based models or large language models (LLMs) however, both of them face challenges. |
| Approach: | They propose a model collaboration pipeline to iteratively optimize a BERT-based corrector. |
| Outcome: | The proposed model outperforms existing methods and outperformed human annotation methods. |
Copied to clipboard
| Challenge: | Existing methods rely on output-level signals for sample identification, such as predictive entropy or semantic similarities with test-time data, which overlook models’ internal dynamics which could pinpoint specific knowledge gaps. |
| Approach: | They propose a Neuron-Aware Active Few-Shot Learning framework that shifts the selection paradigm from output-level proxies to models’ internal dynamics. |
| Outcome: | Experiments on three datasets show that NeuFS outperforms existing AFSL baselines. |
Copied to clipboard
| Challenge: | Existing text editing benchmark datasets contain coarse-grained instructions and lack explainability, resulting in outputs that deviate from intended changes. |
| Approach: | They propose a benchmark specifically designed for fine-grained instruction-based explainable text editing. |
| Outcome: | The proposed benchmark incorporates fine-grained instructions and gold-standard edit explanations. |
Copied to clipboard
| Challenge: | Using gender identity-based framing, language–gender associations are often grounded in the author’s gender identity, inferred from their language use. |
| Approach: | They propose to operationalize the language–gender association as a perceived gender expression of language, focusing on how expression is externally interpreted by humans, independent of the author’s gender identity. |
| Outcome: | The first dataset of itskind identifies 5,100 human annotations of perceived gendered style—human-written texts rated on a five-point scale from very feminine to very masculine. |
Copied to clipboard
| Challenge: | Video-language models excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity. |
| Approach: | They propose a framework that trains video-LLMs to distinguish accurate representations from carefully crafted adversarial examples. |
| Outcome: | Experiments show that VideoPASTA improves performance without human annotation or captioning . the framework can be used on various state-of-the-art video-LLMs with no human annotation . |
Copied to clipboard
| Challenge: | Existing preference-based reward modeling methods face a recursive dependency where each verifier requires a meta-verifier, leading to continuous and costly dependence on human annotation. |
| Approach: | They propose a dual RM that couples discriminative and generative reward models under a non-parametric meta-reward. |
| Outcome: | The proposed model achieves strong performance across major preference benchmarks and even when trained exclusively on language modality, it exhibits robust cross-modal transfer on Omni-RewardBench. |
Copied to clipboard
| Challenge: | Existing instruction following datasets lack logical coherence across turns, narrow topical breadth and heavy manual effort. |
| Approach: | They propose a pipeline that leverages LLMs’ reasoning capabilities to assemble rich, topic-related single-instruction data into multi-turn dialogues and produce chains that are logically coherent, progressively deepen in content, and span diverse domains without fixed templates or extensive human annotation. |
| Outcome: | The proposed pipeline improves the performance of existing LLMs by integrating multiple topic-related data into multi-turn dialogues without fixed templates or extensive human annotation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are evolving rapidly on code generation tasks. |
| Approach: | They propose to automate the vulnerability code benchmark creation with iterative auto validation. |
| Outcome: | The proposed benchmark covers 232 CWE categories across C/C++, Java, and Python languages. |
Copied to clipboard
| Challenge: | a new study examines the accuracy of Wikipedia's factual inconsistencies . a corpus-level inconsistent detection system can help editors identify inconsistances . |
| Approach: | They propose a corpus-level inconsistency detection system that combines LLM reasoning with retrieval to detect and contextualize potential contradictions for human review. |
| Outcome: | The proposed system can detect inconsistencies in Wikipedia and human review. |